Event-Driven Interoperability
Instead of consumers polling for changes or sources pushing to each consumer, sources publish events — "a patient was admitted", "a result is available", "a notifiable condition was diagnosed" — and any number of consumers subscribe.
It is the right pattern when many parties need to react to the same clinical occurrence, and the wrong one when it is used as the sole record of what happened.
1. Problem
Several systems need to know when something happens: surveillance needs notifiable diagnoses, the shared health record needs new encounters, the referral system needs discharges, the reminder service needs new pregnancies.
Point-to-point notification means the source must know every consumer, and adding a consumer means changing the source — which in health usually means a vendor change request and a release cycle.
2. Context
Applies when: multiple consumers care about the same occurrences; timeliness matters; and consumers change more often than sources do.
3. Architecture
Producers Broker Consumers
───────── ────── ─────────
EMR ────────┐ ┌──▶ Shared health record
Lab ────────┼──▶ ┌────────────────────────┐ ───────┼──▶ Surveillance
CHW app ────┤ │ Topics │ ├──▶ Reminder service
Pharmacy ───┘ │ encounter.created │ ├──▶ Analytics pipeline
│ result.available │ └──▶ Audit / archive
│ condition.notifiable │
│ patient.merged │
└────────────────────────┘
retained, replayable log
Producers publish; they do not know who consumes. Consumers subscribe; they do not affect the producer. New consumers are added without touching anything upstream — which is the whole point.
Events usually flow through or alongside the interoperability layer, which remains responsible for identity resolution and terminology translation before publication. Publishing raw, unresolved events pushes that work onto every consumer.
4. Components
| Component | Role |
|---|---|
| Broker | Kafka, RabbitMQ, NATS, or a cloud equivalent. Kafka's retained log makes replay straightforward, which matters here. |
| Schema registry | Versioned event schemas, so consumers do not break on producer changes |
| Producers | Point-of-service systems, or adapters in front of them |
| Consumers | Independently deployed, independently scaled |
| Dead-letter queue | Events that cannot be processed, with an owner |
| Reconciliation job | Periodic full comparison against the source — the safety net |
5. Event design
Thin or thick?
- Thin (notification): "Encounter 123 was created for patient 456." The consumer fetches details, under its own authorisation. Small payloads, no clinical data on the bus, each consumer sees only what it may see.
- Thick (event-carried state): the full resource travels in the event. Fewer round trips, but clinical data is now sitting in a broker's retained log, and every consumer receives everything regardless of what it is entitled to see.
Prefer thin for clinical data. The privacy properties are much better, the authorisation model stays intact, and the broker does not become a shadow clinical repository requiring its own retention and access governance.
Event content, regardless:
- Event type and schema version
- Event identifier (for idempotency)
- Occurrence time and publication time, as distinct fields
- Subject reference, using the shared identifier
- Source system, for provenance
- Correlation identifier, so a chain can be traced end to end
Version from the first event. Adding a required field to an event schema breaks every consumer; a versioning scheme agreed in advance avoids a coordinated release across five teams.
6. The guarantees, and what to do about them
This is where event-driven architectures fail in health, and the failures are quiet.
| Reality | Consequence | Design response |
|---|---|---|
| At-least-once delivery | The same event may arrive twice | Every consumer must be idempotent — use the event identifier |
| No global ordering | An update may arrive before the create | Use resource versions and occurrence times, never arrival order |
| Events can be lost | Silent gaps | Periodic reconciliation against the source — mandatory, not optional |
| Consumers fall behind | Data is hours stale while everything looks healthy | Alert on consumer lag and queue age, not only depth |
| Poison messages | One bad event blocks a partition | Dead-letter queue with a named owner and a review process |
| Bursts | Campaigns and bulk imports produce storms | Rate limiting, back-pressure, and consumers that can be scaled |
The reconciliation requirement is the one to insist on. A surveillance system whose case count depends on never missing a message will be wrong, and nobody will notice until an audit. Run a periodic full comparison — daily, or per reporting period — and alert on discrepancies. The reconciliation job is part of the pattern, not an optional extra.
7. Advantages
- Producers and consumers are decoupled; new consumers cost nothing upstream
- Natural fit for surveillance, care coordination, and index maintenance
- Consumers can be scaled and deployed independently
- A retained log supports replay: a new consumer can process history, and a broken consumer can reprocess after a fix
- Bursts are absorbed rather than propagated
8. Disadvantages
- Eventual consistency — a consumer's view lags, and clinicians may see different data in different systems
- Debugging spans systems; you need correlation identifiers and distributed tracing
- The broker is critical infrastructure requiring real operational capacity
- Schema evolution requires governance
- Personal data in a retained log has retention and access implications that are easy to overlook
9. When not to use
- When you cannot make consumers idempotent. Duplicate processing of a clinical event can mean a duplicate order or a double-counted case.
- When the team cannot operate a broker. Kafka in particular is not low-maintenance, and an unmonitored broker fails silently.
- When there are two consumers and no plans for more. A direct call is simpler and easier to reason about.
- When strong consistency is required. A prescription and its dispensing record should not be eventually consistent.
- As the sole record of truth. Events are notifications. The source system remains authoritative.
10. FHIR Subscriptions and brokers
FHIR Subscriptions provide standards-based notification from a FHIR server, and they are the right choice at the boundary of the ecosystem — for an external participant that should not be given broker access.
Inside the ecosystem, a proper message broker is better: it offers ordering within a partition, retention and replay, consumer groups, back-pressure and observability that Subscriptions do not. The common architecture uses both:
External participants ──FHIR Subscription──▶ IOL ──▶ broker ──▶ internal consumers
Do not use FHIR Subscriptions as the internal event bus. The failure modes listed above all apply, with fewer tools to manage them.
Example technologies
| Role | Options |
|---|---|
| Broker | Apache Kafka, RabbitMQ, NATS JetStream, Redpanda, cloud pub/sub services |
| Schema registry | Confluent Schema Registry, Apicurio |
| Stream processing | Kafka Streams, Flink, ksqlDB |
| CDC from source databases | Debezium |
| Standards-based edge notification | FHIR Subscriptions, HL7 v2 messaging |
References
- FHIR Subscriptions Backport IG — https://hl7.org/fhir/uv/subscriptions-backport/
- Apache Kafka — https://kafka.apache.org/
- NATS — https://nats.io/
- Debezium — https://debezium.io/
- Hohpe & Woolf, Enterprise Integration Patterns — https://www.enterpriseintegrationpatterns.com/